Back

Frontiers in Genetics

Frontiers Media SA

Preprints posted in the last 90 days, ranked by how well they match Frontiers in Genetics's content profile, based on 230 papers previously published here. The average preprint has a 0.18% match score for this journal, so anything above that is already an above-average fit.

1
Evaluating the Impact of Principal Component and Mixed Model Approaches on Polygenic Risk Score Portability to Diverse Ancestries in the UK Biobank

Harikrishnan, A. S.; Kelly, C. M.

2026-08-19 genetic and genomic medicine 10.64898/2026.08.17.26360388 medRxiv
Top 0.1%
13.0%
Show abstract

Polygenic risk scores (PRS) offer considerable potential for precision medicine. How ever, their predictive performance often attenuates when applied to populations that differ from the genome-wide association study (GWAS) training population. There are many potential sources of this portability problem, and one relatively under-explored contributor is the presence of residual confounding in GWAS summary statistics. In particular, confounding specific to the training population may contribute to predictive performance that does not transfer to other populations, such that improved control of population stratification could potentially improve PRS portability. Here, we investigated whether varying levels of population stratification adjustment, through the inclusion of principal components and the use of mixed models, altered PRS portability in three broad ancestry groups in the UK Biobank. The PRS were built using European training data for coronary artery disease and type 2 diabetes and subsequently evaluated in South Asian, African, and Latin American participants. We found that increasing PC adjustment did not produce a consistent trend in portability across ancestry groups or phenotypes, despite modest reductions in the LDSC intercept. However, substantial ancestry- and phenotype-specific effects on transferability were observed. Mixed-model association provided no significant change in PRS discrimination or portability. These findings highlight the need for a better understanding of the nature of residual confounding in PRS and whether improving the causal validity of GWAS results can ultimately improve the transferability of predictive accuracy between populations.

2
Uncovering High-Order Epistatic Interactions in GWAS via a Machine Learning-Based Feature Engineering Framework

Byun, J.; Saha, D.; Han, Y.; Shaw, V. R.; Siminovitch, K.; Amos, C. I.

2026-08-09 genomics 10.64898/2026.08.03.742638 medRxiv
Top 0.1%
11.8%
Show abstract

BackgroundGenome-wide association studies (GWAS) often fail to identify higher-order epistatic interactions that contribute to complex inheritance patterns of traits and diseases. While machine learning (ML) can capture non-linear relationships, extracting interpretable insights from these models remains a challenge. We propose a novel tree-based feature engineering framework that uses Classification and Regression Trees (CART) to explicitly encode high-order interaction decision paths as dummy variables. We investigate three path-based encoding strategies: (i) all decision paths, (ii) leaf-node paths only, and (iii) internal-node paths only. This approach aims to transform complex decision boundaries into discrete features that capture nonlinear interactions that are not readily captured by traditional association models. ResultsThe framework was evaluated using genetic data for ANCA-associated vasculitis (AAV). To manage the high dimensionality of the engineered feature space, we applied a comprehensive suite of ML methods across three tasks: (1) Ensemble Learning (Random Forest, XGBoost, and Gradient Boosting Machine); (2) Decision Tree Analysis (CART); and (3) Regression and Classification Tasks (Regularized Linear Regression/LASSO, Support Vector Machine, and Logistic Regression). Stepwise feature selection and regularization were employed to isolate the most informative interaction patterns. Results indicate that incorporating CART-derived interaction paths--particularly those from high-impact regions of the tree--significantly improves classification accuracy and model interpretability compared to using the original feature space alone. ConclusionsThe proposed framework provides a robust, scalable methodology for identifying high-order genetic interactions. By bridging the gap between the predictive power of ensemble ML and the necessity for mechanistic insight, this approach offers a clearer mapping of the combinatorial genetic processes underlying complex diseases. While applied here to AAV, the method is highly adaptable for exploring the genetic architecture of diverse populations and complex traits.

3
Nitrogen use efficiency in pigs is associated with transcriptomic signatures related to amino acid metabolism, immune activity, and nutrient partitioning

Monney, B.; Ewaoluwagbemiga, E. O.; Kasper, C.

2026-07-01 genomics 10.64898/2026.06.26.733976 medRxiv
Top 0.1%
11.7%
Show abstract

Dietary protein restriction challenges the allocation of amino acids to growth and other physiological functions and therefore requires coordinated metabolic adaptation. Domestic pigs provide an informative system in which to study such responses, because nitrogen retention directly affects lean growth and can be quantified accurately under controlled feeding and housing conditions. Under reduced-protein diets, pigs differ in how effectively they retain nitrogen, and this variation has a genetic basis, making them well suited to investigate the molecular regulation of nitrogen use efficiency (NUE). Here, we characterise differential gene expression and enriched pathways in liver and skeletal muscle of more than 80 pigs with two divergent NUE phenotypes (high and low) maintained under the same protein-reduced, ad libitum dietary conditions. The two NUE phenotypes were clearly distinct at the transcriptomic level, with 177 differentially expressed genes in the liver and 133 in the muscle. In the liver, differential expression and enrichment analyses indicate reduced amino acid catabolism, lower inflammatory and detoxification activity, and a metabolic state that favours lipid processing and insulin-related regulation over the use of amino acids as energy sources. In skeletal muscle, they point to reduced lipid uptake, lower reliance on amino acid oxidation, and a greater emphasis on protein synthesis, translational regulation, mitochondrial energy metabolism, and growth-related processes. These gene-level patterns were supported and extended by pathway and gene-set enrichment analyses. Together, the results suggest that high and low-NUE pigs differ through coordinated, tissue-specific molecular adaptations. Overall, variation in NUE appears to reflect coordinated, tissue-specific differences in how nutrients are allocated between energy use, storage, and lean tissue growth.

4
An endogenous retrovirus insertion disrupting bovine ALKBH8 causes a failure-to-thrive syndrome with immunodeficiency associated with juvenile mortality in Brown Swiss cattle

Glatthard, S.; Kadri, N. K.; Seefried, F. R.; Voitl, L. R.; Weber, B. A.; Schwarzenbacher, H.; Meister, S. L.; Gurtner, C.; OGrady, J. F.; Osbahr, M.; Leonard, A. S.; Meylan, M.; Pausch, H.; Droegemueller, C.; Jacinto, J.

2026-07-10 genomics 10.64898/2026.07.09.737535 medRxiv
Top 0.1%
10.0%
Show abstract

The Brown Swiss (BS) cattle breed is one of the major Swiss dairy breeds. Intensive selection and the widespread use of few elite sires in artificial insemination have increased inbreeding and the occurrence of deleterious recessive alleles in the homozygous state. Analyzing life trajectories in large, genotyped cohorts can identify hidden recessive disorders that are difficult to detect using traditional case-control association testing. Long-read DNA sequencing enables precise detection of causal alleles, including structural variants. This study aimed to (1) identify cryptic recessive loci affecting rearing performance in Swiss BS cattle, (2) evaluate their impact on survival, (3) characterize the associated phenotype, (4) identify the causal variant using long-read whole-genome sequencing, and (5) assess its functional impact. Using Homozygous Haplotype Enrichment/Depletion (HHED) mapping, we identified a risk haplotype (BH39) on chromosome 15 spanning from 16,276,819 bp to 16,446,984 bp that was associated with increased juvenile mortality within the first 180 days of life when present in the homozygous state. The BH39 occurred at a frequency of approximately 4.5% in Swiss BS cattle and 5.3% in German and Austrian BS cattle, and homozygous carriers exhibited a significantly reduced first-year survival rate. Five females homozygous for BH39 underwent clinical examination. They all showed recurrent respiratory disease, impaired growth, poor body condition, rough hair coat, and brown-discolored teeth. Pathological examination revealed bronchopneumonia and eosinophilic enteritis. Clinicopathological findings indicated failure to thrive and immunodeficiency. Long-read WGS of two BH39 homozygous calves revealed a private homozygous coding variant that was in high linkage disequilibrium with BH39. The identified structural variant was an insertion of a large transposable element (10.4 kb ERVK[2-1-LTR]) into the third exon of ALKBH8 (NM_001080341.2 c.267_268indel). Full-length RNA sequencing of cerebellum and liver from a homozygous calf revealed that the endogenous retrovirus (ERV) insertion introduces a cryptic transcription termination signal, truncating ALKBH8 mRNA. This study demonstrates that exploring population-scale genomic data and mining thousands of life-history records, followed by veterinary follow-up evaluations and molecular genetic analyses, provides an effective strategy for identifying cryptic recessive disorders that shorten the lifespan of cattle. The findings provide strong evidence that the ERV insertion into the coding sequence of ALKBH8 represents a loss-of-function variant that causes a previously undescribed recessive disorder that results in increased rearing loss. Interpretive summaryWe identified a recessive disorder in Brown Swiss cattle that causes retarded growth, recurrent infections, immunodeficiency, and increased mortality during the first year of life. Using population-scale genomic data, clinical investigations, and long-read sequencing, we linked the disorder to an exonic transposable element insertion disrupting ALKBH8. The identification of the causal variant now enables direct genetic testing and the implementation of genome-based mating strategies to avoid carrier-by-carrier matings and, consequently, prevent the birth of affected homozygous offspring. We demonstrate the utility of integrating large-scale breeding records, veterinary phenotyping, and advanced genomics to identify hidden defects affecting livestock health and productivity.

5
Germline genomic and methylomic dynamics following three generations of early-life metabolic challenges

de Anca Prado, V.; Pertille, F.; Andersson, D.; Mourin-Fernandez, M.; Godia, M.; Jimenez-Chillaron, J. C.; Ruegg, J.; Guerrero-Bosagna, C.

2026-07-24 genomics 10.64898/2026.07.21.739755 medRxiv
Top 0.1%
9.7%
Show abstract

Environmental and dietary factors can exert multigenerational effects on health and development. In this study, we investigated whether early-life metabolic challenge affects the germline genome and epigenome across three generations. Using a murine model of early life obesity via litter size reduction (overnutrition group, ON) and a control group (CT), we followed the paternal lineage focusing on germline genomic and methylation changes employing Genotyping-by-Sequencing (GBS) coupled with methyl-immunoprecipitation (GBS-MeDIP). We found that unrelated ON families clustered together based on identified Single-Nucleotide Polymorphism (SNP), suggesting that the treatment may have genomic impact. Copy number variations (CNVs) events were identified in ON individuals, being enriched in Long Interspersed Nuclear Elements (LINEs) and Long Terminal Repeats (LTRs). While Principal Component Analysis (PCA) of the methylome showed no clear treatment effect, pathway enrichment and regional analyses revealed methylation changes associated with transposable elements and developmental genes. Notably, the ON group exhibited a disruption in the methylation of Repetitive Elements (RE), which was significant in the same type of RE that were also enriched in the observed CNVs. The ON also showed reduced emergence of novel SNPs in offspring compared to the CT group. These findings suggest that multigenerational metabolic challenge can constrain genetic variability and induce genome instability, potentially mediated by transposable element activity rather than widespread changes in DNA methylation. This work highlights the importance of studying both genome and epigenome dynamics under realistic, multigenerational exposure scenarios and suggests that early metabolic challenges can have long-lasting impacts on genomic architecture and evolutionary potential.

6
Genome-Wide Selection Signatures in Nili-Ravi Buffalo (Bubalus bubalis) Reveal a T-Cell Costimulatory and Cytokine-Signaling Gene Network Distinct from Classical Bovine Tuberculosis Candidate Genes

Ahmad, A.; bakar, A.; Laeeque, S. M.; Khan, W. A.; Kaul, H.; Manan, A.; mustafa, h.

2026-08-11 genomics 10.64898/2026.08.10.743898 medRxiv
Top 0.2%
8.0%
Show abstract

Genomic signatures of selection can reveal loci underlying adaptation and disease resistance in livestock populations, but such analyses in water buffalo (Bubalus bubalis) have historically been constrained by the absence of a chromosome-level, species-native reference genome for SNP array data. We re-analyzed genotype data from 85 Nili-Ravi buffalo (Axiom Buffalo Genotyping 90K array, originally positioned using bovine (Bos taurus, UMD3.1) proxy coordinates, by performing a full coordinate liftover to the buffalo-native UOA_WB_1 assembly using an independently published SNP remapping resource. Following quality control (51,209 markers retained), haplotype phasing, and genome-wide integrated haplotype score (iHS) and Wrights Fst (case/control) selection scans, we evaluated 14 classical bovine-tuberculosis (bTB) candidate genes and identified six additional genes with putative immune function through an unbiased genome-wide screen. None of the 14 classical candidates (including SLC11A1, the Toll-like receptors, and IFNG) reached genome-wide significance in either scan. In contrast, six novel loci TNFSF18, IL2RB, TNFRSF19, IRF2, IL15, and CD28 showed significant iHS or Fst signals, four of which (TNFSF18, IL2RB, IL15, CD28) converge functionally on T-cell costimulation and cytokine receptor signaling (KEGG pathways map04660 and map04060, Bos taurus proxy annotation). Using extended haplotype homozygosity (EHH) decay, haplotype furcation structure, and per-marker haplotype counts as three independent lines of corroborating evidence, we classified these six genes into confidence tiers: TNFSF18 and IL2RB showed the strongest, most balanced support, while CD28 and IL15 signals were driven by very few haplotypes (3 and 5 of 30, respectively) and should be interpreted cautiously pending replication. These findings suggest that adaptive, cell-mediated immune signaling rather than the innate/macrophage-centred mechanisms emphasized by existing bTB candidate gene panels may be a more productive avenue for future selection studies in Nili-Ravi buffalo, while underscoring the value of buffalo-native coordinate systems for accurate genomic inference in this species.

7
Gene regulatory co-expression networks decipher potential lncRNA-miRNA-mRNA interactions modulating transcription regulation in neurodegeneration

Venkatesan, A.; Sinha, P.; Basak, J.; Bahadur, R.

2026-07-08 bioinformatics 10.64898/2026.07.03.736295 medRxiv
Top 0.2%
8.0%
Show abstract

Neurodegenerative diseases are complex disorders characterised by progressive neuronal loss and widespread transcriptomic dysregulation; however, the coordinated interactions among coding and non-coding RNAs that contribute to disease progression remain incompletely understood. In this study, RNA-seq datasets from disease-relevant neuronal populations and brain regions representing Alzheimer's disease (AD), Parkinson's disease (PD) and amyotrophic lateral sclerosis (ALS) were analysed using an integrative network-based framework. Differential expression analysis coupled with weighted gene co-expression network analysis identified modules significantly correlated with disease and prioritised highly connected hub genes. Integration of these hub genes with curated RNA interaction database enabled the construction of candidate lncRNA-miRNA-mRNA regulatory networks. Functional enrichment analysis revealed Gene Ontology biological processes associated with synaptic signalling, mitochondrial function, RNA metabolism and neuroinflammatory responses across neurodegenerative conditions. The inferred regulatory networks suggested both disease-specific and shared post-transcriptional regulatory modules involving key hub genes and non-coding RNAs. Additionally, putative sequence variants were identified within untranslated regions of selected hub genes, suggesting potential alterations in miRNA-mediated regulations. Therefore, this study provides a systems-level view of transcriptomic dysregulation across major neurodegenerative diseases and identifies candidate regulatory interactions and molecular targets for future functional investigation

8
Genetic Modeling of Dyadic Behavioral Traits: Implications for Estimation and Interpretation of Variance Components

Jiang, X.; Siegford, J.; Steibel, J. P.

2026-06-12 genetics 10.64898/2026.06.10.731434 medRxiv
Top 0.2%
7.8%
Show abstract

Studying the genomic control of dyadic social interactions is gaining traction in animal genetics. However, genetic modeling of social interactions poses several challenges, one of which is whether social interactions should be treated as dyadic traits or as aggregated traits at the individual level. In this study, we systematically compared two approaches: dyadic models using dyadic traits and marginal models using marginally aggregated traits and we derived the algebraic relationships between their variance components. In the application, we used a published dataset on post-mixing aggression in pigs, including both directed and undirected aggression records collected during the 9-hour period after mixing among 797 finishing pigs in 59 social groups, as an example to show how model choice can affect variance estimation. Results showed that dyadic models can estimate genetic effects and permanent environmental effects by exploiting repeated dyadic interaction records, thereby enabling a more complete understanding of the sources of variation underlying social interactions. In contrast, marginal models can bias the estimation and interpretation of genetic components, as the aggregated genetic variance may be confounded with other variance components due to the aggregation of dyadic traits. Marginal models may also lead to overestimation of social group and residual variance. These results can provide useful guidance for choosing appropriate modeling strategies for social interaction traits.

9
Insights into the genetic architecture of resistance to viral haemorrhagic septicaemia virus in rainbow trout from a genome-wide association study to in vitro CRISPR-Cas9 functional evaluation

Thomas, V.; Collet, B.; Quillet, E.; Marchand, M.; Huetz, F.; Boudinot, P.; Phocas, F.; Lallias, D.

2026-06-11 genetics 10.64898/2026.06.09.731144 medRxiv
Top 0.2%
7.3%
Show abstract

Viral haemorrhagic septicaemia (VHS) is a severe disease affecting rainbow trout (Oncorhynchus mykiss) and a wide range of wild freshwater and marine fish species. VHSV threatens rainbow trout aquaculture, as it may cause 100% mortality in fry. Previous studies identified a quantitative trait locus (QTL) on chromosome 3 associated with resistance to VHSV waterborne challenge and reduced viral replication in fin explants, although these findings were obtained using limited genetic diversity. The objective of this study was to validate and extend the identification of genomic regions associated with resistance to VHSV in the genetically diverse rainbow trout line designated "synthetic." A genome-wide association study (GWAS) was conducted using whole-genome sequences from parents of progeny classified as resistant or susceptible to a VHSV waterborne challenge. While the QTL on chromosome 3 was not validated in the synthetic line, four novel suggestive SNPs associated with survival following VHSV waterborne challenge were identified on chromosomes 6, 8, 17, and 32. Notably, one SNP on chromosome 17 was located within a gene potentially involved in antiviral defence, a paralog of lrp1 (low-density lipoprotein receptor-related protein 1). To further investigate its role, lrp1 function was analysed in vitro using CRISPR-Cas9 genome editing. Three independent lrp1-/- CHSE-EC cell lines were generated and challenged with VHSV. The results showed that lrp1 is not essential for viral entry but may modulate the inflammatory response during VHSV infection in epithelial cell lines.

10
Using Natural Vector Method for Population Genomic Analysis on Human Mitochondrial Genome Data

Guan, M.; Wu, Q.; Zhao, X.; Yau, S. S.-T.

2026-07-16 genetics 10.64898/2026.07.11.737899 medRxiv
Top 0.2%
7.3%
Show abstract

The natural vector method is an important method for the analysis of biological sequences. In this study, we applied this method to population genetic analysis, with the core purpose of using it to evaluate the characteristics of a set of sequences rather than just pairwise comparison. We used the mitochondrial genome dataset from the human 1000 Genomes Project as a dataset to verify the feasibility of this improved natural vector method. The results showed that the modified natural vector method could be used for various population genetic approaches at least in the sense of population average, including the calculation of principal component analysis, population structure analysis and genetic diversity parameters. The results were in good agreement with those based on traditional molecular genetic markers such as SNP. The new method validates the feasibility of natural vector method for population genetic analysis and provides a framework for the application of matchless pair method to population genomic analysis on a wider scale.

11
Genomic insights into bacterial kidney disease resistance in Arctic charr (Salvelinus alpinus) via a 72k SNP array

Palaiokostas, C.; Jeuthe, H.; Nilsson, K. N.; Hallbom, H.; Axen, C.; Evensen, O.; Eriksson, S.; Johnsson, M.

2026-06-27 genetics 10.64898/2026.06.25.734482 medRxiv
Top 0.2%
7.3%
Show abstract

Selection for disease resistance forms one of the most highlighted areas of aquaculture breeding. A breeding program for Arctic charr has been operating in Sweden for over 40 years, making it the oldest of its kind worldwide for this species. However, the lack of available genomic resources prevented selection for any disease-resistance traits. A 72k Axiom SNP array was produced in this study and used to assess the potential to select for charr resistant to bacterial kidney disease (BKD), which is currently a major threat to the industry. Following a challenge experiment with Renibacterium salmoninarum, the causative agent of BKD, relevant phenotypic proxies were collected from approximately 2,000 charr. Thereafter, those animals were genotyped with the new 72k SNP array. The magnitude of the estimated variance components suggested potential for breeding for BKD resistance in charr, with relevant heritabilities ranging from 0.05 to 0.56 depending on the resistance proxy used. In addition, GWAS suggested that BKD resistance is a polygenic trait. Furthermore, genomic prediction approaches indicated that BKD-resistant animals can be identified using their SNP genotypes. Accuracies, expressed as Pearson correlation coefficients, when BKD resistance was analysed as a continuous trait, ranged from 0.42 to 0.52. In the scenario where BKD resistance was treated as a binary trait, the efficiency of genomic prediction was assessed using ROC curves, with an area under the curve of 0.72. Finally, no unfavourable correlations were found with growth traits. The developed 72k SNP array has the potential of being a pivotal tool for the Swedish Arctic charr breeding program. Moreover, our data support the use of genomic prediction in breeding BKD-resistant Arctic charr. As a critical next step, further validations in actual industry conditions would be required.

12
Blood-based transcriptomic classification of lung cancer: a leakage-free nested cross-validation framework with LASSO

Bakim, S.; UrluOzalan, N.; Gulbahce Mutlu, E.; Demir, V.; Gulbahce, E.

2026-07-13 oncology 10.64898/2026.07.11.26357823 medRxiv
Top 0.2%
7.1%
Show abstract

Peripheral whole-blood gene expression profiling offers a minimally invasive route to lung cancer detection, but high-dimensional transcriptomic data are prone to optimistic bias when preprocessing and model selection are not properly separated from performance evaluation. We applied L1-penalised (LASSO) logistic regression to 303 peripheral whole-blood microarray profiles (123 lung cancer cases and 180 healthy controls; Gene Expression Omnibus accession GSE252168; Illumina HumanHT-12 v4) within a leakage-free nested cross-validation framework (5 outer and 3 inner folds), in which all data-dependent steps (imputation, univariate feature screening by ANOVA F-test with k = 500, and standardisation) were confined strictly to training partitions. Statistical significance was assessed by permutation testing (B = 100), and feature selection stability was quantified across outer folds. LASSO was compared with ridge logistic regression, linear support vector machines, and random forest under the same framework. The LASSO model identified a sparse 29-probe signature with a pooled out-of-fold area under the ROC curve (AUC) of 0.990 (nested estimate 0.989 +/- 0.015), accuracy 97.4%, sensitivity 94.3%, and specificity 99.4% at a 0.50 threshold; permutation testing confirmed significance (p = 0.0099). Six probes, including CDC42, U2AF1, and RPS15A, were selected in all five outer folds, forming a stable core, and all classifiers exceeded AUC 0.987, indicating a strong, algorithm-independent signal. A leakage-free nested cross-validation framework enables unbiased performance estimation and reproducible feature selection in blood-based lung cancer classification. The 29-probe panel is an internally validated candidate requiring prospective, multicentre external validation before clinical use.

13
Assessing the Role of Marker Density and Minor Allele Frequency on Machine Learning Driven Genomic Selection Accuracy in Grapevine

Francisco, F. R.; de Oliveira, G. L.; Niederauer, G. F.; Fritsche-Neto, R.; Souza, A. P. d.; Furlan, M. F. M.

2026-07-17 genetics 10.64898/2026.07.11.737951 medRxiv
Top 0.2%
6.9%
Show abstract

Although grapevine (Vitis spp.) is among the oldest and most economically significant fruit species globally, its genetic improvement faces major bottlenecks due to long juvenile periods and extended cycles for phenotypic evaluation. In this context, genomic selection (GS) has emerged as an effective alternative to traditional selection, offering a robust framework to optimize breeding programs by significantly reducing generation intervals while enhancing predictive accuracy (PA) in early generations and expected genetic gains (EGGs). Nevertheless, factors such as minor allele frequency (MAF) and population size can significantly affect predictive models, even to the point of making their use unfeasible in breeding programs. In this context, this study evaluated the effect of data dimensionality reduction on GS accuracy by selecting single-nucleotide polymorphisms (SNPs) based on MAF thresholds. The experimental design tested the predictive capacities of four machine learning (ML) algorithms (ElasticNet, K-Neighbors, Support Vector Machine Regression, and XGBoost) alongside the conventional Genomic Best Linear Unbiased Prediction (gBLUP) model. These were validated using three SNP datasets (11,115, 9,494, and 6,100 markers) filtered by MAF levels of 0.05, 0.1, and 0.2 across six genetic traits, and EGGs were compared between conventional breeding and GS via the breeders equation. The results revealed that the ML models exhibited remarkable stability, with no significant differences in PA across the different MAF-based SNP densities, except for berry length, which showed a substantial difference with XGBoost at an MAF of 0.2. Conversely, gBLUP demonstrated high sensitivity to dimensionality reduction, with its performance significantly impacted by MAF filtering across all the traits. These results suggest that compared with traditional GS models that rely on a genomic kinship matrix, ML-based approaches offer greater flexibility in feature reduction. Additionally, compared with chemical traits, morphological traits generally had greater predictive ability. Furthermore, every GS model provided estimated genetic gains superior to traditional breeding, with improvements ranging from an 8.90-fold increase in berry length to a 2.86-fold increase in total soluble solids, confirming that GS integration is promising for enhancing breeding efficiency in grapevines.

14
A highly penetrant LMNA R541C variant associated with dilated cardiomyopathy leads to dysregulation in metabolism and proliferation pathways in stem cell-derived cardiomyocytes

Keller, T. E.; Koehring, C.; Higgins, B. R.; Yang, J.; Siddiqui, F. A.; Farsaei, F.; Kim, K.; McDonald, T. V.

2026-07-23 genomics 10.64898/2026.07.20.739542 medRxiv
Top 0.2%
6.9%
Show abstract

BackgroundLMNA codes a widely expressed nuclear cytoskeletal protein (lamin A/C) with multiple important functions. Pathogenic LMNA genetic variation may lead to autosomal dominant cardiomyopathy, though the severity and rate of progression can vary with the specific nucleotide change and location. Prior studies showed that induced pluripotent stem cells (iPSC)-derived cardiomyocytes (iCMs) with LMNA R541C exhibited reduced LMNA protein abundance, increased sarcomere disorganization, and abnormal electrophysiology. MethodsWe investigated the LMNA-R541C variant that exhibits a highly penetrant and severe clinical cardiomyopathy phenotype using transcriptomic analysis of iCMs. Patient-derived iPSCs with CRISPR-corrected (clustered regularly interspersed short palindromic repeats) isogenic control cells and CRISPR knock-in LMNA-R541C heterozygous iPSCs were generated for isogenic controlled experiments. ResultsIn differential gene expression analyses we observed that LMNAR541C/WT iPSC-derived cardiomyocytes had consistent perturbations in 123 genes across CRISPR-corrected and knock-in experiments compared to controls. Pathway analysis identified that the G2M checkpoint and oxidative phosphorylation processes were consistently dysregulated and confirm these findings in previously published iPSC and murine models. DiscussionThese results implicate perturbed gene expression and pathways that may contribute to the severe phenotypes in LMNA-R541C. Informatic analysis of pathways suggests several drug classes including multiple cardiac glycosides as potential targeted therapeutic candidates to be explored.

15
Genetic Architecture and Sample Size Impact Relative Performance of Nonlinear Machine Learning and Standard Polygenic Risk Scores

Zhu, J.; Baousi, A.; Morris, A. P.; Guo, H.

2026-09-03 genetic and genomic medicine 10.64898/2026.08.29.26361109 medRxiv
Top 0.3%
6.6%
Show abstract

Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study (GWAS) summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs directly from individual-level data, with the aim of improving predictive performance over standard PRSs through their ability to model non-additive genetic effects. However, their superiority across studies has been inconsistent, and the conditions under which they provide meaningful improvements remain unclear. We combined theoretical analysis, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretical analysis showed that standard PRSs can implicitly capture part of the genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects, thereby losing less information than commonly assumed. Although nonlinear models have a higher theoretical potential, their greater flexibility incurs a bias-variance trade-off that can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed the standard PRS only when the genetic architecture involves a sufficiently large proportion of interaction genetic variance concentrated across relatively few interaction effects and large training samples were available. Random forest consistently underperformed the standard PRS. In an application to ischemic heart disease prediction using UK Biobank data, XGBoost showed no meaningful improvement in predictive performance over the standard PRS, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.

16
Benchmarking Imputation Methods for Single-Cell RNA Sequencing Data Using Peripheral Blood Mononuclear Cells from Acute Myocardial Infarction Patients

Ramesh, P.; Fyta, M.

2026-08-27 bioinformatics 10.64898/2026.08.23.746230 medRxiv
Top 0.3%
6.5%
Show abstract

Acute myocardial infarction (AMI) remains one of the leading causes of mortality worldwide, and the following post-effects, such as post-AMI inflammation and tissue repair, involve peripheral blood mononuclear cells playing a critical role. The influence of imputation methods in biological data is assessed with respect to high-resolution single-cell RNA sequencing (scRNAseq) data relevant to these cells. Still scRNAseq data often encounter a lot of dropout events, leading to sparse and noisy datasets, hampering downstream results. To assess the influence of the missingness in the data, we artificially impose different levels of dropout in available scRNAseq data by leveraging various imputation techniques. Specifically, we introduce artificial missingness at 10%, 20%, and 30% levels under a missing completely at random (MCAR) framework, repeated across 10 independent runs. We benchmarked six imputation strategies - MAGIC, IterativeImputer, KNNImputer, Mean Imputation, SoftImpute, and a Generative adversarial network (GAN) - based approaches using multiple evaluation metrics: marker gene preservation, clustering consistency (Adjusted Rand Index - ARI), gene-wise correlation with ground truth, and structural separation (silhouette scores). The results clearly underline that no single imputation method dominated across all metrics. Overall, Mean and KNN imputers showed limited recovery across all benchmarks. GAN excelled in global transcriptional recovery and SoftImpute in preserving biologically meaningful cell-type signals. Our results highlight the importance of selecting the imputation methods as part of the pre-processing step towards the downstream biological questions related to transcriptome recovery, detection of marker genes, or maintaining cell-type-specific resolution.

17
Two-tower models for genomic prediction of reproductive outcomes and sex-specific fertility liabilities: simulation insights

Pappas, F.; Palaiokostas, C.; Debes, P. V.; Johnsson, M.

2026-07-09 genetics 10.64898/2026.07.03.736358 medRxiv
Top 0.3%
6.1%
Show abstract

Many biological characteristics arise by interactions between more than one biological organism or unit. Fertilization success in sexually reproducing species represents such an extended phenotype where both mates are required to be fertile for a successful outcome. Consequently, predictive models should account for the joint nature of reproductive performance while offering interpretable estimates for individual mate contributions. Recent advances in genomics and machine learning (ML) provide standardized, high-dimensional genetic information on one hand and computational tools capable of modeling complex biological systems on the other. Here, we construct and evaluate two-tower (TT) machine learning architectures for genomic prediction of binary reproductive outcomes and recovery of sex-specific fertility liabilities. Simulated datasets, generated under a range of genetic architectures, were utilized to compare multilayer perceptron (TT-MLP), convolutional neural network (TT-CNN), and L1-regularized linear (TT-LASSO) two-tower models. Simulation scenarios varied sex-specific heritabilities, genetic correlations, infertility prevalence, mating structure, and sex-specific infertility rates. Models were evaluated with regard to their ability to predict reproductive success at pair level and also recover true underlying genetic values for male and female fertility. Prediction accuracy increased with the underlying heritable component as expected, while sex-specific tower-scores successfully recovered latent fertility liabilities despite models being trained only on observed joint outcomes. TT-LASSO achieved the highest overall classification performance, whereas TT-MLP provided more balanced and consistent recovery of sex-specific genetic values across scenarios. An additional simulation, incorporating genotype-dependent mate compatibility demonstrated advantages of fully-connected neural networks for capturing non-additive interactions. These results indicate that two-tower frameworks provide a powerful approach for modeling reproductive traits, enabling simultaneous prediction of aggregate reproductive outcomes and sex-specific fertility liabilities from genotypic information.

18
Long-term realized genetic gain and population dynamics under genomic selection in Brazilian cassava germplasm

de Freitas, G. M.; Certuche, D. C. S.; Jannink, J.-L.; De Oliveira, E. J.; Garcia, A. A. F.

2026-07-22 genetics 10.64898/2026.07.18.739356 medRxiv
Top 0.4%
5.6%
Show abstract

Genomic selection has become an important strategy in cassava breeding, enabling faster selection cycles and sustained genetic progress. Despite its widespread adoption, long-term evaluations integrating predictive performance, realized genetic gain, and genetic diversity remain scarce, particularly in clonally propagated crops. We present a comprehensive assessment of genomic selection outcomes in the Brazilian cassava breeding program across four recurrent selection cycles (C0 to C3) implemented between 2011 and 2024, using historical phenotypic and genomic data from 210 multi-environment trials. Predictive ability of genomic best linear unbiased prediction models ranged from low to moderate, depending on the traits genetic architecture and heritability. Prediction accuracies were highest in early cycles (C0 and C1) and showed modest declines in later cycles (C2 and C3). Root yield, shoot yield, plant height, starch content, and dry matter content exhibited stable predictive performance across cycles, with a gradual reduction in RMSE, indicating improved model calibration as training populations expanded. Regression analyses of genomic estimated breeding values revealed significant realized genetic gains for most yield-related traits. In contrast, dry matter content and starch content exhibited small, non-significant negative trends, consistent with known unfavorable genetic correlations with yield. Targeted reductions in plant architecture scores reflected deliberate selection for ideotypes suited to mechanized production systems. At the same time, analyses of genetic diversity revealed a slight decrease in observed heterozygosity, with higher values in the most advanced selection cycle. These results provide an integrated framework for monitoring predictive performance, realized genetic gain, and population genetic dynamics under long-term genomic selection. Collectively, they offer valuable insights into balancing short-term genetic improvement with long-term sustainability and support the development of strategies to optimize selection decisions, breeding planning, and population management in Brazilian cassava breeding programs.

19
Functional Characterization of Transcriptome-Wide Isoform Switching in Hürthle Cell Carcinoma (HCC)

Butt, R. S.; Amir, A.; Paracha, R. Z.

2026-07-27 bioinformatics 10.64898/2026.07.23.740299 medRxiv
Top 0.4%
5.6%
Show abstract

Hurthle cell carcinoma (HCC) is an aggressive form of thyroid cancer. While mitochondrial DNA mutations and chromosomal losses have been identified in HCC, isoform switching, and its functional consequences remain uncharacterized. This study reanalyzed NCBI GEO dataset GSE228870 (n = 32), using Salmon and IsoformSwitchAnalyzeR() to identify isoform switching. The analysis resulted in 371 switches across 335 genes showing functional consequences including loss of protein domains, shorter open reading frames (ORFs), loss of signal peptides and novel sub-cellular localizations. Most significant isoform switches (q-value < 0.05, |dIF| > 0.1) were observed in LAMA2, LSP1, MAD2L2, FBLN2 and CXCL12, implicating extracellular matrix dysregulation, DNA damage response, immune signaling and cytoskeleton regulation. These genes are expressed in normal thyroid (median TPM 20.69, 11.66, 14.79, 134.1 & 80.76). However, specific isoforms of LAMA2 and MAD2L2 are not expressed in normal thyroid, explaining tumor-specific expression in HCC. Alternative transcription termination site (ATTS) gain was significant, suggesting altered 3 end in HCC transcripts. TCGA SpliceSeq showed LSP1, FBLN2 and CXCL12 undergo alternative promoter (LSP1 exon1 PSI=94.5%, FBLN2 exon2 PSI=99.0%) and alternative termination (CXCL12 exon3.3 PSI=53.9%) in thyroid cancer, suggesting ATTS and alternative transcription start site (ATSS) as shared splicing dysregulation mechanisms. This is the first systematic characterization of isoform-level dysregulation in HCC.

20
Machine Learning-based Prediction of Preterm Birth Using Genetic Data

Sundelin, H.; Jacobsson, B.; Ytterberg, K.; Sole-Navais, P.; Juodakis, J.

2026-06-26 genetic and genomic medicine 10.64898/2026.06.24.26356330 medRxiv
Top 0.4%
5.6%
Show abstract

The leading cause of mortality and morbidity in children under the age of 5 is preterm birth. The timing of birth is influenced by both genetic and environmental factors, but the underlying mechanisms remain poorly understood, making its prediction difficult. In this study, we investigated the potential of using machine learning models to predict preterm birth based on genetic data from the Norwegian Mother, Father and Child Cohort Study (MoBa). We trained and evaluated several classification algorithms on individual-level genetic data from over 15,000 mothers and children. Our results indicate that the predictive capacity of maternal gestational duration-associated loci for preterm birth is limited, with the highest AUC values around 0.57. Additionally, incorporating more SNPs within the associated loci did not improve prediction performance. As expected, the contribution of the maternal genome to preterm birth prediction was found to be larger than that of the fetal genome. Overall, our findings suggest that while genetic testing provides some information about an individual's risk for preterm birth, further research incorporating additional factors is necessary to enhance predictability.